Papers with proxy metric

2 papers
RoSE: Round-robin Synthetic Data Evaluation for Selecting LLM Generators without Human Test Sets (2026.eacl-long)

Copied to clipboard

Challenge: Current large language models (LLMs) are powerful generators of synthetic data, which are used for training smaller, more efficient models.
Approach: They propose a proxy metric for selecting the best LLM generator without human annotations and a metric that measures the performance of a model.
Outcome: The proposed proxy metric outperforms intrinsic heuristics and comes within 0.76 percentage points of the optimal generator baseline.
The Woman Worked as a Babysitter: On Biases in Language Generation (D19-1)

Copied to clipboard

Challenge: a systematic study of biases in natural language generation (NLG) is presented . a study of language models in NLG is conducted by examining language models.
Approach: They propose a systematic study of biases in natural language generation by analyzing text generated from prompts that contain mentions of different demographic groups.
Outcome: The proposed method reveals biases in natural language generation (NLG) by analyzing text generated from demographic prompts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations